Papers with fine-grained evaluation framework
TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing evidence-based summarization tasks require tracing source evidence to assess their accuracy. |
| Approach: | They propose a benchmark for traceable, aspect-based summarization that pairs summaries with sentence-level citations to enable users to trace back to the original context. |
| Outcome: | The proposed benchmark can be used to evaluate document summarization with LLMs and human evaluations. |
EXAMS: A Multi-subject High School Examinations Dataset for Cross-lingual and Multilingual Question Answering (2020.emnlp-main)
Copied to clipboard
| Challenge: | EXAMS is a benchmark dataset for cross-lingual and multilingual question answering for high school examinations. |
| Approach: | They propose to use EXAMS to evaluate cross-lingual and multilingual question answering for high school examinations. |
| Outcome: | The proposed model can be used to explore multilingual reasoning and knowledge transfer methods and pre-trained models in schools in different languages, which was not possible by now. |
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to transfer text style focus on sentence-level data, limiting performance . current LLMs struggle to generate public speaking texts that align with human preferences . |
| Approach: | They propose a task to transform official texts into public-speaking styles by analyzing real-world data. |
| Outcome: | The proposed task aims to transform public speaking texts into public-speaking styles . the proposed framework analyzes characteristics and identifies problems of stylized texts . |
FinGrAct: A Framework for FINe-GRrained Evaluation of ACTionability in Explainable Automatic Fact-Checking (2025.findings-emnlp)
Copied to clipboard
| Challenge: | despite the importance of actionability, no prior research has evaluated its effectiveness. |
| Approach: | They propose a fine-grained evaluation framework that can access the web to assess actionability in AFC explanations. |
| Outcome: | The proposed framework surpasses state-of-the-art evaluators in achieving highest correlation with human judgments while showing lowest egocentricbias. |
A Unified Agentic Framework for Evaluating Conditional Image Generation (2025.acl-long)
Copied to clipboard
Jifang Wang, Yangxue Yangxue, Longyue Wang, Zhenran Xu, Yiyu Wang, Yaowei Wang, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang
| Challenge: | Conditional image generation is a popular and personalization-oriented task, but there are challenges in developing task-agnostic, reliable, and explainable evaluation metrics. |
| Approach: | They propose a unified agentic framework for comprehensive evaluation of conditional image generation tasks. |
| Outcome: | The proposed framework achieves a high correlation with human assessments on seven prominent image generation tasks. |
Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation Framework (2023.acl-long)
Copied to clipboard
| Challenge: | Current evaluations of FEC models that depend on factuality metrics are not reliable and detailed enough. |
| Approach: | They propose a fine-grained evaluation framework that automatically evaluates FEC models on different error categories. |
| Outcome: | The proposed evaluation framework compares models on different error categories and finds the best training modes and significant differences in the performance of existing models. |
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study (2025.findings-emnlp)
Copied to clipboard
Yujun Zhou, Jiayi Ye, Zipeng Ling, Yufei Han, Yue Huang, Haomin Zhuang, Zhenwen Liang, Kehan Guo, Taicheng Guo, Xiangqi Wang, Xiangliang Zhang
| Challenge: | Existing benchmarks that rely on final-answer accuracy fail to capture the quality of the reasoning process. |
| Approach: | They propose a fine-grained evaluation framework that assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
| Outcome: | The proposed framework assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
BAGELS: Benchmarking the Automated Generation and Extraction of Limitations from Scholarly Text (2025.findings-emnlp)
Copied to clipboard
| Challenge: | a growing number of scientific publications have limitations as a source of uncertainty. |
| Approach: | They propose a computational architecture for extracting and generating limitations from scholarly papers using a novel Retrieval Augmented Generation technique. |
| Outcome: | The proposed architecture extracts limitations from ACL, NeurIPS, and PeerJ papers and supplementes them with external reviews. |
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies on table-based reasoning focus on a single gold table, not multiple tables . a persistent demand for robust table understanding systems is resulting from the complexity of table data . |
| Approach: | They propose a MT-RAIG Bench to evaluate systems on Retrieval-Augmented Insight Generation over Mulit-Tables. |
| Outcome: | The proposed framework improves human quality judgments on the generated insights. |
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing models fail to recall and accurately apply designated persona knowledge without explicit cues . memory-driven role-playing paradigms are attracting significant interest . |
| Approach: | They propose a memory-driven role-playing paradigm that frames persona knowledge as the LLM's internal memory store and a prompting architecture that guides structured memory retrieval and response generation. |
| Outcome: | The proposed paradigm provides a comprehensive diagnostic for four-stage role-playing abilities across 12 LLMs. |
Understanding GUI Agent Localization Biases through Logit Sharpness (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Multimodal large language models often exhibit hallucinations that compromise reliability . despite promising performance, these models often display systematic localization errors . |
| Approach: | They propose a framework that categorizes model predictions into four distinct types . they propose metric that evaluates alignment between semantic continuity and logits distribution . |
| Outcome: | The proposed framework categorizes model predictions into four different types . it reveals nuanced failure modes beyond traditional accuracy metrics . |
CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLM (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing medical NLP benchmarks focus on qualitative reasoning and textual comprehension, but lack of fine-grained evaluation of intermediate reasoning. |
| Approach: | They propose a Chinese medical calculation benchmark that disentangles clinical entity extraction from numerical computation. |
| Outcome: | The proposed framework disentangles clinical entity extraction from numerical computation, enabling systematic diagnosis of model deficiencies. |